Genome Medicine
○ Springer Science and Business Media LLC
Preprints posted in the last 7 days, ranked by how well they match Genome Medicine's content profile, based on 183 papers previously published here. The average preprint has a 0.18% match score for this journal, so anything above that is already an above-average fit.
SULAIMAN, M. A.; Oyeyemi, B. F.
Show abstract
Sub-Saharan African populations carry pharmacogenomic alleles poorly represented in the European-derived reference panels underlying most clinical genotyping tools. We present a curated, machine-readable catalog of nine actionable alleles across six pharmacogenes (CYP2D6, CYP2B6, CYP2C9, CYP2C19, CYP3A5, NAT2) with African-specific frequency ranges, functional annotations, and evidence levels derived from reanalysis of 661 high-coverage whole-genome sequences across seven 1000 Genomes Project African populations. Direct comparison against PharmCAT v3.4.0 shows that CYP2D6 produces zero diplotype calls (0/661 samples callable) due to monomorphic reference positions absent from standard variant-only VCF output, a known limitation whose consequences for African allele carriers had not been reported. afripharmagen's reduced-position strategy identifies 243 CYP2D617 and 134 CYP2D629 carriers from the same input. For CYP2B6, CYP2C9, CYP2C19, and NAT2, both tools show concordance of 95-100%. Frequency gradients (CYP2B66: 30-50%; CYP2D617: 15-35% in West Africa; CYP3A5*1: 60-95%) translate directly into prescribing risk for efavirenz, tramadol, tacrolimus, and isoniazid. Pharmacogenomic decision support in African settings must incorporate population-specific allele definitions and input-format-aware strategies.
Hasan, A.; Demidova, E. V.; Priyadarshini, P.; Czyzewicz, P.; Gathuka, L.; Murayama, T.; Zhou, Y.; Kiss, Z. A.; Shastry, R. K.; Andrake, M.; Hearne, G.; Devarajan, K.; Wu, C.; Shah, A.; Schultz, B. M.; Connolly, D. C.; Rosen, G. L.; Canadas, I.; Liu, J. C.; Burtness, B. A.; Smith, J. J.; Dunbrack, R. L.; Golemis, E. A.; Whetstine, J. R.; Meyer, J. E.; Arora, S.
Show abstract
Chemoradiotherapy (CRT) is the standard-of-care therapy for many solid malignancies, yet predictive biomarkers of treatment response remain limited. We identified a germline single nucleotide polymorphism (SNP) in an intrinsically disordered region of the lysine demethylase KDM3C/JMJD1C (p.S464T) that is associated with CRT outcomes in locally advanced rectal cancers (LARC) and head and neck squamous cell carcinoma (LA-HNSCC). In silico modeling with AlphaFold predicted S464T substitution influenced interaction between phosphorylated KDM3C and RNF8 FHA domain. In cellular models, conversion of S464 to T464 increased sensitivity to DNA-damaging agents. S464T substitution impaired damage-induced MDC1-RAP80 signaling and downstream RAP80-BRCA1 colocalization. SNP carrying cells impaired DNA repair causing genotoxic stress that is associated with increased cGAS-cGAMP innate immune signaling and increased apoptosis. Population analyses with the SNP highlighted an increase incidence of UV-induced skin and other cancers, linking inherited variation in the chromatin regulatory gene KDM3C to genome instability, cancer risk, and therapeutic vulnerability.
Yarmolinsky, J.; Cavallo, F. R.; Koskeridis, F.; Yu, X.; Bouras, E.; Richenberg, G.; Costantini, I.; Ray, D.; Woolf, B.; Karhunen, V.; Ellis, L.; Haycock, P. C.; Hemani, G.; Davey Smith, G.; Tsilidis, K. K.; Zuber, V.; McKay, J. D.; Dehghan, A.; Tzoulaki, I.
Show abstract
Confounding is a central challenge in observational studies. Here, we propose a framework for identifying confounders of two non-causally related traits by employing cross-trait pleiotropy analysis to detect genetic loci that affect both traits and multi-trait colocalisation to identify molecular phenotypes mediating these effects. We apply this approach to the analysis of C-reactive protein (CRP) - a non-specific marker of inflammation - and 10 inflammation-related cancers. In UK Biobank, higher pre-diagnostic CRP levels are associated with increased risk of multiple cancers, but bidirectional Mendelian randomization provides little evidence for a causal relationship. Cross-trait genetic analyses identify 92 loci with shared CRP-cancer effects including those with established roles in cancer and 50 novel loci such as RSPO3 (breast cancer) and GCKR (colorectal cancer). Integration with proteomic and single-cell transcriptomic data identified putative molecular mediators at 24 loci including plasma TLR1 levels in breast cancer and CD4+ T cell IRF5 expression in kidney cancer. Notably, 15 candidate effector genes encode targets of approved or investigational medications, including IL6, PDE4D, and CASP8, indicating potential opportunities for their repurposing for cancer prevention. The proposed approach provides a generalisable framework for leveraging non-causal phenotypic relationships to yield insights into disease mechanisms and therapeutic targets for disease prevention.
Luo, X.; Syreeni, A.; Hill, C.; Smyth, L. J.; Dahlstrom, E. H.; Mutter, S.; Chen, Z.; Natarajan, R.; Pan, S.; Parton, A.; Jackson, H.; McKay, G.; Susztak, K.; Hirschhorn, J. N.; Florez, J. C.; Maxwell, A. P.; Groop, P.-H.; McKnight, A. J.; Sandholm, N.
Show abstract
Hyperglycaemia is a hallmark of diabetes and a major risk factor for diabetic kidney disease (DKD). However, the molecular consequences of long-term cumulative hyperglycaemia (CH) remain unclear. As a stable epigenetic modification, DNA methylation may capture past glycaemic exposure. Here, we assessed CH-associated DNA methylation in 1,245 participants with type 1 diabetes (T1D) from Finland and the United Kingdom-Republic of Ireland cohorts. We identified 17 CH-associated CpGs, with the strongest association at cg19693031 (TXNIP). Longitudinal analyses demonstrate that these CH-associated DNA methylation levels remain stable despite short-term glycaemic fluctuations, suggesting lasting epigenetic imprints of earlier metabolic control. Integrative analyses combining genomic, epigenetic, and proteomic data characterized these CpGs and potential target proteins. Mendelian randomization suggested a causal association between cg20853880 (KLF11) and DKD, supported by chromatin accessibility and kidney KLF11 expression. Our findings suggest that epigenetic changes contribute to metabolic memory and may mediate the effects of hyperglycaemia on DKD.
Efthymiou, S.; Tabata, K.; Dafsari, H. S.; Schober, E.; Latza, C.; Isaoglu, M.; Abuelrub, A.; Rad, A.; Firoozfar, Z.; Turchetti, V.; Lin, R. Q.; Maroofian, R.; Wiethoff, S.; Afzal, E.; Zafar, F.; Rana, N.; McRae, A. M.; Kaiyrzhanov, R.; Guliyeva, U.; Gulieva, S.; Melikishvili, G.; Lespinasse, J.; Vitobello, A.; Denomme-Pichon, A.-S.; Wentzensen, I. M.; Mefford, H. C.; Briere, L. C.; A Walker, M.; A High, F.; Sweetser, D. A.; Kendall, M.; Franchi, M.; Brown, M.; Latner, D.; Joset, P.; Ivanovski, I.; Alfadhel, M.; Alluhaydan, I.; Frederiksen, A. S.; Arriens, V.; Hanker, B.; Mankad, K.; Guerin, J
Show abstract
Pathogenic variants in RUBCN, encoding the Run domain Beclin-1 interacting and cysteine-rich domain-containing protein (Rubicon) have been implicated in autosomal recessive spinocerebellar ataxia 15 (SCAR15). However, the molecular mechanisms underlying disease pathogenesis remain poorly understood. Here, we report 18 individuals from 15 unrelated families harbouring biallelic RUBCN variants, who present with an aggressive neurodevelopmental disorder variably characterized by seizures, developmental delay, intellectual disability and movement abnormalities that cause regression, progressive brain atrophy and neurodegenerative features. Through functional characterization, we demonstrate that a subset of disease-associated putative truncating variants disrupt autophagy regulation. In Caenorhabditis elegans models, loss-of-function RUBCN variants result in an increased autophagic flux and impaired neuronal function, recapitulating key features in humans. Correspondingly, cellular assays reveal that nonsense and frameshift RUBCN variants lead to defective autophagy inhibition, underscoring a crucial role for RUBCN as a key negative autophagy regulator. Molecular dynamics simulations rank the eleven missense variants by structural effect, with p.Arg813Trp alone altering the target protein at both the local and the regional level and lying within the RAB7A-binding module that the truncating alleles remove altogether. Our findings establish and expand the RUBCN-related disorders as a clinically and molecularly distinct subset of autophagy-related diseases. By delineating both the genetic landscape and cellular consequences of Rubicon dysfunction, this study enhances our understanding of autophagy-related neurodevelopmental disorders and provides a foundation for future therapeutic investigations.
Zhao, L.; Zeng, Y.; Abelman, D. D.; Lin, W.; Luo, P.
Show abstract
Motivation: Cell-free DNA methylation provides a minimally invasive signal for early cancer detection and tissue-of-origin prediction. Most methods represent methylation measurements as independent fixed-window features and therefore do not explicitly model relationships among genomic regions. Results: We developed PANGEM (Pan-cancer Graph-based Cancer Detection Using the Cell-free DNA Methylome), a graph-learning framework that represents genomic bins as nodes and integrates CpG context, genomic proximity, and sample-specific methylation similarity in the graph topology. Across five repeated stratified train-test splits, PANGEM achieved the highest mean performance among evaluated methods, with an AUROC/AUPR of 0.997/1.000 for binary cancer detection and macro-AUROC/AUPR of 0.977/0.870 for multiclass tissue-of-origin prediction. In the independent INSPIRE cohort, 72 of 78 cancer cases (92.3%) exceeded the binary classification threshold, and PANGEM correctly classified 9 of 17 head and neck cancer cases (52.9%), the highest accuracy among evaluated methods. Subnetwork analysis further identified recurrent, graph-connected methylation patterns, including a 111-DMR subnetwork with increased methylation in cancer samples.
Bresnahan, S. T.; Xiong, C.; Head, T.; Chang, Y.-H.; Bhattacharya, A.; Huang, J. Y.
Show abstract
Unmeasured confounding threatens causal inference and replicability in observational multi-omic studies across variable environments. Genetic instrumental variables (Mendelian randomization) and negative-control calibration each address complementary sources of unmeasured confounding, yet no existing framework unifies them for omics-scale mediation analysis. We introduce ICONIC, an R package that embeds genetic instruments and negative controls within a proximal causal inference framework for total-effect and mediation analysis. ICONIC implements eight estimators spanning five confounding-control strategies, supports continuous, binary, and time-to-event outcomes, and provides extensive diagnostics including sensitivity analyses that map estimator performance across plausible assumptions. Ground-truth benchmarks are calibrated to real-omics covariance structures via a hybrid generative model (GAN + feature-level Gaussian copula) rather than parametric simulation, and a companion planning tool predicts performance gains from collecting additional omic data. We demonstrate ICONIC in two case studies: identifying placental transcriptomic mediators of gestational diabetes on birth weight (n = 164), and tumor-expression mediators of smoking intensity on lung cancer survival (n = 494). Notably, ICONIC's diagnostics recommended different estimation strategies across the two scenarios, reflecting differences in the likely influence of unmeasured confounding. ICONIC is freely available at https://github.com/sbresnahan/iconic/.
Ivankovic, F.; Ko, A.; Aster, M. M.; Balaconis, M. K.; Banks, E.; Bemis, M.; Cibulskis, K. R.; Degatano, K.; Gauthier, L. D.; Grant, G.; Hatcher, A.; Kachulis, C.; Karczewski, K. J.; Labrecque, S. M.; Lawson, J.; Liao, C.; Magner, R.; Munshi, R.; Schatz, M. C.; Schultz, P. M.; Shah, S. P.; Sheets, E. A.; Tibbetts, K.; Vernest, K. A.; Ye, R.; Gabriel, S.; Lennon, N. J.; Neale, B. M.; Browning, B. L.; Lichtenstein, L. T.
Show abstract
Genotype imputation remains essential for large-scale human genetics studies, but its performance is limited by the size and ancestral diversity of available reference panels, reducing accuracy for rare variants and underrepresented populations. Here, we present a cloud-based imputation service built on a multi-ancestry reference panel derived from 515,579 jointly phased genomes from the All of Us (N=414,830) and National Human Genome Research Institute's Analysis, Visualization, and Informatics Lab-space (AnVIL, N=100,749) datasets. The All of Us + AnVIL reference panel is highly diverse and includes 261,163 participants most genetically similar to non-European reference populations, spanning 665,398,839 high-quality autosomal sites, representing a nearly 50% increase over TOPMed, the previous largest imputation service. Across multiple ancestry groups, the panel enables accurate imputation (empirical R2 0.8) for variants with allele frequencies as low as 0.2%, extending reliable imputation into the rare-variant frequency spectrum, including allele frequencies down to 0.002% and 0.006% for samples with European ancestry and African ancestry in the United States, respectively. Compared with TOPMed, the panel improves imputation accuracy across all ancestry groups except Africans, and recovers additional trait-associated variants not represented in existing reference panels. To facilitate broad community access while preserving participant privacy, we deploy the panel through a secure cloud-based imputation platform using privacy-preserving recombined haplotypes. This resource establishes a new foundation for genome-wide association studies (GWAS) and fine-mapping, especially in previously underrepresented populations.
Li, D.; Feng, Q.; Zhang, Y.; Chen, H.; Wang, X.; Shen, C.
Show abstract
Background National childhood respiratory pathogen spectra are diversifying nearly everywhere - within-country diversity rose in 203 of 204 countries between 1990 and 2023 - yet whether countries are diversifying toward a common spectrum or along divergent paths is unknown. We quantified between-country compositional distance of national pathogen spectra over the same period. Methods We built national pathogen share vectors from Global Burden of Disease Study 2023 lower respiratory infection etiologic attributions (26 pathogens, 204 countries, ages 0-19 years) at five timepoints spanning 1990-2023. Between-country distance was measured as all pairwise Jensen-Shannon divergences (JSD; primary) and Bray-Curtis dissimilarities, with Baselga and Jaccard decompositions; robustness was assessed across metrics, pathogen panels, low-count thresholds and a balanced panel of 107 countries. Results Mean pairwise JSD rose from 0.0084 in 1990 to 0.0283 in 2023 (+238%; trend p = 0.030), peaking in 2021 (+283%) with a partial 2023 pullback. Bray-Curtis dissimilarity rose +120% and the balanced panel +423%. Divergence was entirely balanced variation (share reallocation), with spectrum richness rising from 18.5 to 21.1 of 26 pathogens. Dispersion rose fastest for influenza (coefficient of variation 0.03 to 0.55) and respiratory syncytial virus (0.08 to 0.48). Within-region distance rose in every computable GBD super-region (five of seven): divergence occurs within regions, not between blocs. Conclusions National spectra are re-sorting along country-specific axes as vaccine-preventable dominance recedes at different speeds. Diversification is universal, but convergence is absent: the transition at the etiologic-spectrum level is asynchronous and path-dependent, with implications for empirical treatment policy and pathogen surveillance.
Weyrich, M.; Ware, A.; Steixner-Kumar, A.; Windschmitt, J.; Sarakpi, T.; Abplanalp, W.; Dimmeler, S.; Speer, T.; Zeiher, A. M.
Show abstract
Clonal hematopoiesis (CH) increases with age, but whether different somatic clones represent an ageing phenotype or exert distinct systemic effects is unclear. In 450,587 UK Biobank participants, including 46,324 with plasma proteomics, we compared clonal hematopoiesis of indeterminate potential (CHIP) and mosaic loss of chromosome Y (mLOY) or X (mLOX) across biological ageing, incident disease, and circulating proteins. Despite shared age dependence, these alterations showed distinct disease spectra: non-DNMT3A CHIP was associated with broad multisystem disease burden, mLOY with a more focused respiratory, musculoskeletal and cardiovascular profile, whereas mLOX lacked broad age-related disease associations. Clone burden mapped to distinct proteomic programs: mLOY to neutrophil degranulation and extracellular-matrix remodeling, non-DNMT3A CHIP to myeloid immune regulation, and mLOX unexpectedly to cytotoxic lymphocyte/NK-cell responses. Mendelian randomization supported selected protein-disease relationships. Thus, age-related hematopoietic clones are not interchangeable markers of ageing but define alteration-specific systemic programs associated with distinct disease vulnerabilities.
Tiwari, P.; Garg, M.; Pattanayak, S.; Sarkar, I.; Roy, R.; Bhatraju, N.; Verma, A.; K, S. R.; Prakash, S.; Kumar, V. S.; Uddin, M. A.; Rawat, N.; Sahu, A.; Kumar, Y.; Leuva, P. H.; Mridha, A.; Yenamandra, V.; Singh, A. P.; Mishra, A.; Raychaudhuri, S.; Tallapaka, K. B.; Chandak, G. R.; Kulkarni, M. J.; Dharne, M.; Wahengbam, R.; Kalita, J.; Manna, P.; Subudhi, U.; Majumder, S.; Chakraborty, P.; Chaudhary, K.; Sengupta, S.; Phenome India Consortium, ; Sardana, V.; Chatterjee, S.; Ganguly, D.
Show abstract
Background: India has a rising incidence of chronic non-communicable diseases, making it a major healthcare burden today. Growing evidence suggests that chronic low-grade inflammation links ageing with cardiometabolic disorders, captured by the emerging concept of inflammaging. However, most evidence on biological ageing comes from Western populations, with no similar models developed for the Indian population. Given the country's distinctive genetic makeup, unique exposome, and heterogeneous NCD presentation, Western models may not capture inflammaging and its effects in the Indian population. Methods: We analysed baseline data from 4,240 adults in the Phenome India CSIR Health Cohort Knowledgebase (PI CheCK), a nationwide multi-centre cohort. Participants were stratified into eight cardiometabolic phenotype groups by BMI (Asian cut off), blood pressure and HbA1c status. We trained a Super Learner ensemble to predict chronological age in the lean normotensive-normoglycaemic reference group (n=615) using 44 plasma cytokines, sex, haemoglobin, and bioimpedance-derived visceral fat area, per cent body fat, and total body water. Performance was assessed by repeated five-fold cross-validation and in a held-out healthy test set. Calibrated biological age acceleration was then estimated in the remaining 3,625 participants. Results: Median age was 51.0 years (IQR 41.0 to 62.0) and 49.4% were female. The Super Learner outperformed elastic net and XGBoost comparators. Permutation importance identified visceral fat area, per cent body fat, CTACK, SDF1a, haemoglobin and sex as leading contributors, with body composition measures accounting for the largest share, indicating an immune-metabolic rather than cytokine-only signal. Biological age acceleration was concentrated in overweight/obese phenotypes. Lean phenotypes showed acceleration close to the reference (0.32 0.50 years). Conclusions: Cytokine and body composition measures capture a quantifiable immunometabolic ageing signal in a South Asian cohort, with acceleration driven predominantly by adiposity. External validation and longitudinal follow up are required.
Chia, C.; Baker, K.
Show abstract
Obesity is a significant public health concern. Early-onset obesity in the context of rare disease can reflect genetically-mediated pathology or elevated susceptibility through indirect mechanisms. Mapping the diverse characteristics and needs of young people with obesity in the rare disease population is a first step toward mechanistic and translational research. We carried out a retrospective comparative analysis of demographic, genotypic, phenotypic and health service utilisation data for young people with obesity (cases: n=500) and without obesity (controls: n=11,444) from the UK 100,000 Genomes Project rare disease cohort. Cases and controls were recruited prior to genomic diagnosis, across clinical disorder categories. We observed significant association between socioeconomic deprivation and obesity risk. Young people with obesity had significantly higher utilisations of acute care and mental health services, indicating an overall higher health burden. A curated panel of 519 candidate obesity-associated genes demonstrated aggregate association with obesity, although no single gene reached significance. Phenotypic comparison between cases and controls highlighted increased multi-organ and neurological system involvement, highlighting the overlap between neurodevelopmental and obesity risks. Within the case group, we conducted cluster analysis to identify early-onset obesity groups with different phenotypic profiles, potentially arising from different causal pathways - this identified six obesity subgroups of interest, with differing involvement of neurodevelopmental and other systems. Our study confirms that obesity co-occurs with a wide range of factors within the rare disease population, and is associated with significant physical and mental health needs, requiring holistic lifelong care.
Venkatesh, R.; Deo, R.; Cappola, T.; Penn Medicine BioBank, ; Ritchie, M. D.; Kim, D.
Show abstract
Atrial fibrillation (AF) is the most common sustained cardiac arrhythmia and a major cause of cardioembolic stroke. Although polygenic risk scores (PRS) are well characterized to quantify inherited susceptibility for AF, they provide limited insight into the pathways and tissues underlying genetic risk, which are critical to uncover for individual risk prediction. In this study, we develop a pathway-level multi-omics representation learning framework that converts individual genetic profiles into interpretable biological features by integrating GWAS-derived pathway burden scores with tissue-specific transcriptomic pathway signals. We constructed machine learning models to assess population-level AF risk prediction performance across genomic and transcriptomic tissue contexts; the pathway-based global attention models substantially improved risk prediction performance over PRS and other baselines (AUROC improved from 0.601 to 0.738). Transformer and graph neural network frameworks then assessed individual-level pathway interpretability, revealing heterogeneous contributions from electrical signaling, cardiac development, and DNA repair pathways to AF risk. This added interpretability highlights the potential of this pathway approach to enable more mechanistically informed risk stratification than static PRS by capturing underlying heterogeneity. To independently assess whether prioritized pathways reflected cardiac regulatory biology, we compared pathway rankings with transcriptional effects predicted by the AlphaGenome foundation model. Variants in highly ranked pathways showed significantly greater predicted effects on expression in atrial and ventricular tissues (FDR = 0.032) relative to controls, providing orthogonal evidence that the model identifies biologically relevant mechanisms. Overall, this work reframes polygenic risk from a single measure of susceptibility to tissue-informed pathway mechanisms, providing a framework for interpretable genomic stratification in complex diseases.
Hessel, M.; Inda Diaz, J. S.; Sjöberg, A.; Salva-Serra, F.; Helldal, L.; Jirstrand, M.; Johnning, A.; Kristiansson, E.; Skovbjerg, S.
Show abstract
Antimicrobial resistance is a public health challenge, driving the need for rapid, cost-effective diagnostic support tools. Artificial intelligence (AI) may enable prediction of susceptibility to untested antibiotics from known susceptibility results, but prospective clinical validation is required before routine use. We evaluated an AI-based decision support method, trained on invasive isolates from the European Surveillance System (TESSy), for prediction of antibiotic susceptibility in clinical Escherichia coli urine isolates. The evaluation included 99 E. coli isolates from urine samples with diversity in age, sex, and antibiotic susceptibility. Predictions were evaluated for 14 antibiotics using patient metadata and susceptibility results for 4-8 antibiotics as input. Prediction uncertainty was handled using conformal prediction, allowing abstention when confidence was insufficient. EUCAST disk diffusion test results were used as reference and genomic sequence data was used to explore mechanisms of the AI performance. Without conformal prediction, 84% of predictions were correct when susceptibility results of six antibiotics were used to predict susceptibility to eight additional antibiotics. Across all predictions generated using susceptibility results for six antibiotics as input, the major and very major error rates were 19% and 12%, respectively. Prediction errors varied between antibiotics and were associated with certain phenotypic and genotypic resistance patterns. Conformal prediction reduced errors but increased abstentions; at confidence levels of 90%, 95%, and 97.5%, the model abstained in 9.6%, 14%, and 22% of instances. The method showed promising performance, but its clinical use remains limited and may require diagnostic data beyond susceptibility test results and demographic variables.
Bowness, J. S.; Bernal Martinez, A.; Barinka, J.; Schulte-Schrepping, J.; Renders, S.; Waclawiczek, A.; Leppa, A.-M.; Trumpp, A.; Raffel, S.; Haas, S.; Velten, L.
Show abstract
To sustain blood formation, hematopoietic stem and progenitor cells (HSPCs) coordinate a multitude of cell biological processes, from cell cycle control and stress responses to lineage priming. While many genetic regulators of high-level HSPC function have been identified, how HSPCs coordinate more basal cell biological programs, and how such programs relate to stem cell function, remains incompletely understood. Here we use Perturb-seq to profile the transcriptional consequences of targeting 520 genes by CRISPRi in primary mouse HSPC cultures. We developed an analytical strategy to separate perturbation-induced changes in cell-state abundance and clonal heterogeneity from cell-state-local transcriptional effects. From these local perturbation signatures, we identified 19 gene regulatory programs (GRPs) that are defined by co-regulation in response to genetic perturbation, in contrast to co-expression or human curation, and align well with cell biological processes. By decomposing gene expression data from functional and clinical studies into program activity, we show that GRP activities associate with, and predict, phenotypes such as clonal output after transplantation, as well as survival and drug response in retrospective acute myeloid leukemia (AML) cohorts. Together, our study establishes perturbation-derived co-regulation programs as an interpretable framework for linking genetic regulators, cell-biological processes and stem-cell-associated phenotypes.
Yang, Y.; Vasudevaraja, V.; Serrano, J.; Mohamed, H.; Kelly, S.; Jour, G.; Gindin, T.; Park, K.; Jones, D.; Feng, X.; Pinnell, J.; Mclennan, S.; Tin, M. Y.; Tsirigos, A.; Snuderl, M.; Wrzeszczynski, K. O.
Show abstract
Next-generation sequencing (NGS) for the detection of somatic variants has become the method of choice in a variety of molecular oncology fields and in the clinic. Its use ranges from sequencing entire tumor genomes and transcriptomes to targeted clinical diagnostic gene panels. The NYU Langone Genome PACT (Profiling of Actionable Cancer Targets, LG-PACT) assay is a qualitative in vitro diagnostic test that uses targeted next generation sequencing (NGS) of formalin-fixed paraffin-embedded (FFPE) tumor tissue matched with normal specimens from patients to detect gene alterations in a targeted panel covering 606 genes and the TERT promoter. Indications for testing are cancer (solid tumors and hematological malignancies) where a mutational profile from multiple genes would be informative for disease stratification, prognosis, or treatment options including targeted therapies and eligibility for clinical trials. The test is intended to provide information on somatic mutations including point mutations, small insertions/deletions (indels), and copy number aberrations for diagnostic and treatment decisions. LG-PACT is a United States Food and Drug Administration (FDA) cleared diagnostic test (510K: K202304). The clinical interpretation of sequencing data of molecular tumor markers from NGS encompasses automated variant calling tools with human interpretation. This final mostly manual review of data step is intensive, involving highly trained scientists, encompassing literature review, interpretation and clinical tier classification by pathologists, who then provide a complete molecular diagnostic report to the treating oncologists. We provide analysis of 1339 clinical genomic profiles from 31 different cancers and their subtypes, comprising of central nervous system (CNS) 792 (59%) cases (incl. meningioma, glioma and glioblastoma), with 267 (20%) cases predominantly of lung, pancreatic and colorectal and 280 of others (21%). Here, we present the technical challenges of validating an NGS oncological diagnostic targeted assay for clinical grade accuracy and sensitivity for patient care. We show how copy number alterations provide a more comprehensive description of the tumors genomic profile. We then outline the utility of targeted panel sequencing based on certified pathologist selection of reportable variants for our current patient cohort. Where analysis of variant detection has led to 49.4% (661/1339) of our clinical tumor samples containing mutations in known therapy targeted genes, 35.6% (477/1339) with mutation detected in other genes, and 15% (201/1339) cases being negative.
Kouam, C.; Mingle, J.; Alvarez Jerez, P.; Evans, A.; Moller, A.; Baker, B.; Weller, C.; Paquette, K.; Brooks, J.; Grant, S. M.; Ayuketah, A.; Meredith, M.; Palade, J.; Malik, L.; Hise, K.; Raphael Gibbs, J.; Anderson, J.; Ding, J.; Harbert, R.; Fu, Y.; Zheng, X.; Garcia-Ruiz, S.; Gustavsson, E. K.; Blauwendraat, C.; Ryten, M.; Sedlazeck, F.; Ferrucci, L.; Reed, X.; Nalls, M. A.; Cookson, M. R.; Van Keuren-Jensen, K.; Hutchins, E.; Jain, M.; Billingsley, K. J.
Show abstract
Isoform-resolved transcriptomics is fundamental to decoding the molecular complexity of the human brain, yet population-scale long-read RNA sequencing has remained inaccessible due to labor-intensive library preparation, sensitivity to RNA degradation in postmortem tissue, and the absence of integrated, reproducible analysis pipelines. Here we present SALRR (Scalable Analysis of Long-Read RNA-seq), an integrated wet-lab and computational platform designed to overcome these barriers. Automated ONT long-read cDNA library preparation on the Hamilton Microlab NGS STAR platform reduces hands-on time by 67% and enables 24 libraries per operator per day while maintaining performance across RNA integrity values. A modular, Snakemake-based pipeline performs end-to-end processing from ONT signal data to isoform-level quantification, incorporating SIRV spike-in calibration, multi-stage quality control, and stringent isoform validation. Applied to 10 postmortem frontal cortex samples from the North American Brain Expression Consortium, SALRR identified 31,607 high-confidence isoforms from 10,075 genes, including 8,532 novel splice variants absent from GENCODE v49, and complex splicing events systematically missed by short-read sequencing at neurodegeneration-relevant loci, including GBA1, CCNF, CHCHD10, and TREM2. All protocols and code are openly available, providing a scalable, community-ready framework for isoform-resolved transcriptomics in neurodegeneration, aging, and complex brain disease.
Erhart, D. K.; Ressin, H.; Balz, L. T.; Chatterjee, S.; Lule, D.; Mueller, S.; Lewerenz, J.; Muench, J.; Tumani, H.; Gross, R. M.
Show abstract
Post-COVID-19 syndrome (PCS) is characterized by fatigue, neurological impairment and systemic symptoms. This heterogeneity of symptoms hinders biomarker development. Here, we profiled extracellular-vesicle (EV) surface markers in plasma and CSF from 61 participants with PCS (COVIDpost), 80 recovered controls (COVIDreco), and 10 participants with non-SARS-CoV-2 post-viral syndromes. EVs were analysed by bead-based multiplex flow cytometry using tetraspanin-directed (TSPN) and phosphatidylserine-directed lactadherin (PS) detection. Amongst 37 targets covering tetraspanins and vasculature-, immunity- and stemness-associated markers, none met a 1% false-discovery-rate threshold. However, L1-regularized logistic regression under fully nested 5x5 cross-validation identified a distributed plasma EV profile, with mean out-of-fold areas under the receiver operating characteristic curve (AUCs) of 0.788 (95% CI 0.715 - 0.852) for TSPN and 0.716 (95% CI 0.636 - 0.792) for PS detection. Across the pooled COVIDpost and COVIDreco population, EV classification scores covaried with clinical group differences, but did not track clinical severity within either cohort. These PCS-EV classification scores decreased at one-year follow-up in COVIDpost participants. Our findings identify an internally cross-validated multivariable EV surface profile associated with COVIDpost versus COVIDreco status and support independent validation and exploration of EV-based biomarkers in post-viral fatigue syndromes.
Takeuchi, J. S.; Kurokawa, M.; Yamamoto, K.; Yamanaka, J.; Morino, E.; Takayanagi-Nishisako, S.; Ohmagari, N.; Sugiura, W.; Kimura, M.
Show abstract
Background The COVID-19 pandemic substantially altered respiratory pathogen circulation worldwide. However, longitudinal analyses of changes in respiratory pathogen ecology across the pandemic and post-pandemic periods remain limited. Methods We analyzed 19,968 respiratory samples tested with the BioFire(R) FilmArray(R) Respiratory Panel at a hospital in Tokyo, Japan, between January 2020 and March 2026. We evaluated temporal changes in pathogen circulation, age-specific epidemiology, co-detection patterns, pairwise pathogen associations, and clinical parameters. Results At least one respiratory pathogen was detected in 27.8% of tests. Respiratory pathogens resurged asynchronously following the relaxation of COVID-19-related public health measures. Influenza virus circulation remained markedly suppressed until late 2022 before re-emerging in successive large seasonal epidemics, whereas other pathogens, including RSV, human metapneumovirus, and Mycoplasma pneumoniae, exhibited distinct resurgence patterns. Pathogen distributions also varied by age. Human rhinovirus/enterovirus remained predominant among young children, whereas SARS-CoV-2 predominated among older adults. Co-detection occurred in 14.0% of positive specimens and was significantly more frequent in younger patients. Pairwise analysis identified both positive and negative pathogen associations; however, the patterns varied across age groups and study periods. Conclusions Respiratory pathogen circulation changed substantially during the transition from the COVID-19 pandemic to the post-pandemic period, with pathogen-specific, age- and period-dependent patterns. Continued surveillance is warranted to determine how respiratory pathogen circulation will evolve and to inform infection control strategies in the post-pandemic era.
Elena, A. X.; Batantou Mabandza, D.; Kluemper, U.; Breurec, S.; Dagot, C.; Berendonk, T. U.
Show abstract
The global dissemination of antimicrobial resistance is increasingly driven by bacterial clones combining antimicrobial resistance with enhanced virulence and environmental adaptability. Escherichia coli sequence type 131 (ST131) has historically been regarded as a major disseminator of the extended-spectrum {beta}-lactamase (ESBL) blaCTX-M-15. However, the emergence of E. coli ST1193 carrying blaCTX-M-15 may represent an ongoing shift in the epidemiology of this resistance determinant. Here, we investigated the prevalence, genomic characteristics, virulence and antimicrobial resistance potential of ST1193 in comparison with ST131. A total of 1,136 E. coli isolates were recovered from touristic and non-touristic environments, hospital-associated samples, and aircraft toilets in Guadeloupe. Isolates were whole-genome sequenced and analysed for antimicrobial resistance and virulence determinants. Additionally, publicly available genomic data comprising 1,215 blaCTX-M-15-positive ST131 and ST1193 isolates were analysed to assess temporal and geographical trends. ST1193 was significantly associated with aircraft-associated samples and exhibited a higher antimicrobial resistance gene burden than ST131, while maintaining a comparable virulence factor content. Analysis of publicly available genomes revealed similar temporal emergence patterns for blaCTX-M-15-positive ST1193 and ST131, with ST1193 showing a more recent distribution and a higher number of deposited isolates in recent years, consistent with a potential ongoing clonal replacement. Comparative genomic analysis identified numerous virulence and adaptation-associated genes shared between both sequence types, while ST1193 additionally carried distinct determinants, including components of the transmissible locus of stress tolerance. Furthermore, quinolone resistance-associated mutations were strongly linked to blaCTX-M-15 carriage, particularly among ST1193 isolates. Together, these findings identify E. coli ST1193 as an emerging high-risk clone with substantial potential for blaCTX-M-15 dissemination. Its association with aircraft-associated samples further highlights the potential role of air travel in long-distance transmission and underscores the need to reconsider current surveillance strategies focused predominantly on ST131.